Skip to content

P3-8: CSI500 independent generalization check for value/lowvol signals - #21

Merged
StackOverFlow11 merged 3 commits into
mainfrom
p3-csi500-value-lowvol-generalization
Jun 13, 2026
Merged

P3-8: CSI500 independent generalization check for value/lowvol signals#21
StackOverFlow11 merged 3 commits into
mainfrom
p3-csi500-value-lowvol-generalization

Conversation

@StackOverFlow11

Copy link
Copy Markdown
Owner

Summary

P3-8: does the P3-7 sign-level conclusion (value_ep/value_bp positive, volatility_20 negative on the independent 2024-07..2026-05 window) generalize outside the screened universes, specifically to CSI500 (000905.SH, mid-caps) — a cell independent in BOTH universe and time? Report-only; the unchanged P3-7 independent-validation machinery (same factor groups, same cost scenarios, same hypotheses); no new factors, no tuning, no alpha/portfolio/execution/OOS-slicing change.

  • Config config/phase3_real_csi500_generalization.yaml: screened anchor SSE50|2022-2024 (must reproduce P3-6/P3-7), independent anchor SSE50|2024-2026 (must reproduce the P3-7 verdict numbers), and the NEW independent cell 000905.SH|2024-2026; CSI500|2022-2024 skipped + disclosed (runtime budget). Data feasibility probed first (monthly 500-name index_weight snapshots through 2026-05-29 incl. a pre-start 2024-06-28 snapshot).
  • output.subset_report_name (the baseline_report_name precedent): each subset-validation study owns its report filename, so the P3-8 run does not clobber the accepted P3-7 artifact. Configs without the key keep the historical filename bitwise (locked by tests).
  • output.subset_report_title (review HIGH fix): the report H1 is config-driven so the CSI500 study names itself ("Phase 3-8 — CSI500 Independent Generalization Check") instead of the machinery's default P3-7 phase label; the body framing stays sample-aware. P3-6/P3-7 configs leave it unset and keep the renderer's sample-aware default (regression-locked).

Real run (3 cells, ~3.55h; the 735-name CSI500 cell dominates)

  • Both anchors reconciled exactly: screened raw ICs 22/22 vs the P3-5 matrix report; the independent SSE50 verdict ICs identical to P3-7 (a second reproducibility confirmation).
  • CSI500 verdict: SUPPORTED (21 settled rebalances vs minimum 8): value_ep +0.0083/+0.0145, value_bp +0.0230/+0.0127, volatility_20 −0.0350/−0.0272 (train/test, both subperiods hold).
  • Generalizes — and with LESS attenuation than the SSE50/CSI300 holdouts (CSI500 test |IC| 0.013–0.027 vs 0.003–0.016 there): the value/low-vol signs hold outside the screened universes; combo_ic_weighted CSI500 test IC positive in all 4 groups (highest of the three holdout cells).
  • Honest caveats (disclosed): CSI500 portfolio results are positive across all groups and the full cost ladder, but legacy_trio posts +17.8%→+10.6% — another small-sample/regime ranking flip across cells; ~21 rebalances, single window. NOT a return claim.
  • Secret scan: 0 occurrences in report and log; the P3-7 artifact is not overwritten.

Test plan

  • pytest -p no:cacheprovider383 passed (+8: CSI500 config cells/groups/scenarios/hypotheses, sample classes, default + own report filename, config-driven title incl. "first line is NOT Phase 3-7", default-title regression)
  • ruff check . — clean
  • validate-config on all 11 configs — OK
  • run-phase0 demo regression — ic 0.9600 / annual 0.8408 unchanged
  • git diff --check + conflict marker scan — clean
  • AGENTS.md == CLAUDE.md (verbatim copy convention)
  • secret scan on the merge diff — 0 token occurrences, no artifacts/parquet committed
  • real CSI500 run completed; artifact title fixed (Phase 3-8), anchors reconciled, 2/2 independent verdicts SUPPORTED. The 3.5h matrix was NOT rerun for the title fix (deterministic 1-line H1 re-render; real numbers untouched).

…aming)

Ask whether the P3-7 sign-level conclusion (value_ep/value_bp positive,
volatility_20 negative on the independent 2024-07..2026-05 window)
generalizes outside the screened universes, specifically to CSI500
(000905.SH) — a cell independent in BOTH universe and time. Machinery is
the unchanged P3-7 independent-validation layer: same factor groups,
same cost scenarios, no new factors, no tuning, no alpha/portfolio/
execution/OOS-slicing change.

- config/phase3_real_csi500_generalization.yaml: screened anchor
  SSE50|2022-2024 (must reproduce P3-6/P3-7), independent anchor
  SSE50|2024-2026 (must reproduce the P3-7 verdict numbers), and the
  NEW independent cell 000905.SH|2024-2026; CSI500|2022-2024 skipped +
  disclosed (runtime budget). Data feasibility probed: monthly 500-name
  index_weight snapshots through 2026-05-29 incl. a pre-start
  2024-06-28 snapshot; 645 distinct in-window constituents.
- output.subset_report_name (baseline_report_name precedent): each
  subset-validation study owns its report filename; None keeps the
  historical default bitwise (locked by tests) — a P3-8 run no longer
  clobbers the accepted P3-7 artifact.
- tests: +4 (CSI500 config cells/groups/scenarios/hypotheses, sample
  classes, default report filename preserved, P3-8 owns its filename)
  -> 379 passed; ruff clean; 11 configs validate; demo run-phase0
  unchanged (ic 0.96 / annual 0.8408)
- CLAUDE.md / AGENTS.md (verbatim copy): P3-8 progress entry — both
  anchors reconciled exactly (screened raw ICs 22/22 == P3-5; the
  independent SSE50 verdict ICs identical to P3-7, a second
  reproducibility confirmation); CSI500|2024-2026 verdict SUPPORTED
  (value_ep +0.0083/+0.0145, value_bp +0.0230/+0.0127, volatility_20
  -0.0350/-0.0272; 21 settled rebalances vs minimum 8) with LESS
  attenuation than the SSE50/CSI300 holdouts -> the P3-7 sign-level
  conclusion GENERALIZES outside the screened universes; CSI500
  portfolio results positive across all groups and the full cost ladder
  (honest caveat: legacy_trio +17.8% is another small-sample/regime
  ranking flip; not a return claim); gates updated to 379 passed
- TEST_REPORT.md: P3-8 per-file breakdown (4) + real-data validation
  entry (anchor reconciliation, verdict, preserved P3-7 artifact,
  secret scan 0)
- RUNBOOK.md: Phase 3-8 section (cell roles, subset_report_name,
  report-type title note, skip disclosure)
…(review HIGH)

The P3-8 CSI500 report rendered with the hardcoded "Phase 3-7 —
Independent-Sample Validation" H1 (qt/reports.py only branched the title
on has_independent), contradicting the config and docs that identify it
as the P3-8 CSI500 generalization study — the same class of stale report
wording flagged on P3-7.

- config: OutputCfg gains `subset_report_title` (the `subset_report_name`
  precedent — config owns the study identity). None keeps the renderer's
  sample-aware default (P3-7 independent / P3-6 post-hoc).
- render_subset_validation: the H1 is config-driven when
  output.subset_report_title is set, else the existing sample-aware
  default. The body framing (screened-vs-independent mechanics) stays
  sample-aware — correct for any such study, so only the H1 changes.
- config/phase3_real_csi500_generalization.yaml: sets the title to
  "Phase 3-8 — CSI500 Independent Generalization Check ...".
- tests: +4 (config sets title; P3-6/P3-7 leave it unset; rendered
  CSI500 H1 starts with Phase 3-8 + contains CSI500 + is NOT Phase 3-7;
  default title preserved without the override) -> 383 passed.
- artifact: regenerated deterministically — only the H1 line changes
  with this fix, so the accepted real numbers are untouched (the 3.5h
  matrix is NOT rerun); diff is exactly the 1 title line; secret scan 0.
- docs: RUNBOOK/TEST_REPORT/CLAUDE.md/AGENTS.md note the config-driven
  title and the corrected "title carries report type" wording.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant